Papers with interlinear glossing task
Functional Lexicon in Subword Tokenization (2025.naacl-long)
Copied to clipboard
| Challenge: | Function units are hard to map across languages, while being the most frequent tokens. |
| Approach: | They analyze subword tokens in terms of their productivity and try to find thresholds that best distinguish function from content tokens. |
| Outcome: | The proposed method can be used to identify functional lexical units in low-resource languages with minimal annotated data. |